Speaker-Dependent Voice Activity Detection Robust to Background Speech Noise

نویسندگان

  • Shigeki Matsuda
  • Naoya Ito
  • Kosuke Tsujino
  • Hideki Kashioka
  • Shigeki Sagayama
چکیده

In this paper, we proposed a speaker-dependent voice activity detection (VAD) algorithm that only extracts the speech period uttered by a target user. Based on our survey of the recognition error of real speech data collected in “VoiceTra,” which is a speech-to-speech translation system for smartphones, we found that many word insertion errors are caused by background speakers’ speech. Our VAD, which consists of three GMMs (noise and speaker-independent GMMs as used in a traditional GMM-based VAD and a speaker-adapted GMM) can be used for speech detection of the target speaker. In the VAD evaluations for the test utterances with background speakers’ speech, our proposed VADs achieved better performance than the conventional VAD. Also speech recognition experiments demonstrated that an ASR system with our proposed VAD achieved better performance than an ASR system using the conventional VAD.

برای دانلود رایگان متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

A New Algorithm for Voice Activity Detection Based on Wavelet Packets (RESEARCH NOTE)

Speech constitutes much of the communicated information; most other perceived audio signals do not carry nearly as much information. Indeed, much of the non-speech signals maybe classified as ‘noise’ in human communication. The process of separating conversational speech and noise is termed voice activity detection (VAD). This paper describes a new approach to VAD which is based on the Wavelet ...

متن کامل

Robust Voice Activity Detection for Interview Speech in NIST Speaker Recognition Evaluation

The introduction of interview speech in recent NIST Speaker Recognition Evaluations (SREs) has necessitated the development of robust voice activity detectors (VADs) that can work under very low signal-to-noise ratio. This paper highlights the characteristics of interview speech files in NIST SREs and discusses the difficulties of detecting speech/non-speech segments in these files. To alleviat...

متن کامل

Artificial Neural Network-Based Feature Combination for Spatial Voice Activity Detection

For many applications in speech communications and speechbased human-machine interaction, a reliable Voice Activity Detection (VAD) is crucial. Conventional methods for VAD typically differentiate between a target speaker and background noise by exploiting characteristic properties of speech signals. If a target speaker should be distinguished from other speech sources, these conventional conce...

متن کامل

Comparison of Voice Activity Detectors for Interview Speech in NIST Speaker Recognition Evaluation

Interview speech has become an important part of the NIST Speaker Recognition Evaluations (SREs). Unlike telephone speech, interview speech has substantially lower signal-to-noise ratio, which necessitates robust voice activity detection (VAD). This paper highlights the characteristics of interview speech files in NIST SREs and discusses the difficulties in performing speech/nonspeech segmentat...

متن کامل

The QUT-NOISE-SRE protocol for the evaluation of noisy speaker recognition

The QUT-NOISE-SRE protocol is designed to mix the large QUT-NOISE database, consisting of over 10 hours of background noise, collected across 10 unique locations covering 5 common noise scenarios, with commonly used speaker recognition datasets such as Switchboard, Mixer and the speaker recognition evaluation (SRE) datasets provided by NIST. By allowing common, clean, speech corpora to be mixed...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

عنوان ژورنال:

دوره   شماره 

صفحات  -

تاریخ انتشار 2012